Papers with human reasoning
Beyond Multiword Expressions: Processing Idioms and Metaphors (P18-5)
Copied to clipboard
| Challenge: | idioms and metaphors processing is a rapidly growing area in NLP, says dr. s. robertson . idiomatic idiomas are characteristic to all areas of human activity and to all types of discourse. |
| Approach: | This tutorial will provide attendees with a clear notion of idioms and metaphors . it will provide them with computational models of linguistic characteristics and methods . |
| Outcome: | This tutorial aims to provide attendees with a clear notion of the linguistic characteristics of idioms and metaphors . it outlines how to model idiomatic idiomes and their processing and what resources are available to support their use . |
Learning to Imagine: Integrating Counterfactual Thinking in Neural Discrete Reasoning (2022.acl-long)
Copied to clipboard
| Challenge: | Existing NDR models suffer from large performance drop on hypothetical questions, e.g., “what the annualized rate of return would be if the revenue in 2020 was doubled”. |
| Approach: | They propose a learning to imagine module which can be seamlessly incorporated into NDR models to perform the imagination of unseen counterfactual. |
| Outcome: | The proposed model can perform the imagination of unseen counterfactuals on hypothetical questions. |
„Mann“ is to “Donna” as「国王」is to « Reine » Adapting the Analogy Task for Multilingual and Contextual Embeddings (2023.starsem-1)
Copied to clipboard
| Challenge: | a lack of comparable multilingual benchmarks and a consensual evaluation protocol for contextual models remains an open question. |
| Approach: | They propose a multilingual analogy dataset and evaluate human and contextual embedding performance. |
| Outcome: | The proposed dataset evaluates human and contextual embedding models on the analogy task. |
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering (2025.acl-short)
Copied to clipboard
| Challenge: | Current approaches generate visual markers for all questions, generating excessive visual markers. |
| Approach: | They propose a plug-and-play approach that adapts to the complexity of questions . they propose combining fast intuitive judgments with deliberate analytical reasoning . |
| Outcome: | The proposed approach improves performance on four benchmarks on ScienceQA, TextQA, VizWiz, and MME. |
Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason’s Selection Task (2026.eacl-short)
Copied to clipboard
| Challenge: | In humans, reasoning often performs well in domain specific settings, especially in normative rather than purely formal contexts. |
| Approach: | They propose a dataset that explicitly encodes deontic modality to systematically distinguish deontics from descriptive conditionals and analyze LLMs’ conditional reasoning under deontical rules. |
| Outcome: | The proposed dataset systematically distinguishes deontic from descriptive conditionals and examines LLMs’ conditional reasoning under deontics. |
Abstraction-of-Thought Makes Language Models Better Reasoners (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Abstract reasoning is a key to generalization in human reasoning, but eliciting language models to perform reasoning with abstraction remains unexplored. |
| Approach: | They propose a new structured reasoning format called Abstraction-of-Thought (AoT) this approach elicits language models to first contemplate on the abstract level before incorporating concrete details . |
| Outcome: | The proposed model outperforms the prevailing Chain-of-Thought (CoT) reasoning on 23 unseen tasks. |
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback (2026.eacl-long)
Copied to clipboard
| Challenge: | Novelty assessment is a central yet understudied aspect of peer review . manuscript submissions double roughly every 15 years, and individual reviewers now complete an average of 14 reviews per year. |
| Approach: | They propose a structured approach for automated novelty evaluation that models expert reviewer behavior through three stages: content extraction, retrieval and synthesis of related work, and structured comparison for evidence-based assessment. |
| Outcome: | The proposed approach outperforms existing LLM-based baselines on 182 ICLR 2025 submissions with human-annotated reviewer novelty assessments. |
From Detection to Explanation: Effective Learning Strategies for LLMs in Online Abusive Language Research (2025.coling-main)
Copied to clipboard
Chiara Di Bonaventura, Lucia Siciliani, Pierpaolo Basile, Albert Merono Penuela, Barbara McGillivray
| Challenge: | Abusive language detection requires commonsense reasoning, world knowledge and linguistic nuances that evolve over time. |
| Approach: | They propose a knowledge-guided version of Llama-2 instruction fine-tuned for multi-class abusive language detection and explanation generation that mitigates bias and generates explanations that are relevant to the text and coherent with human reasoning. |
| Outcome: | The proposed model mitigates bias and generates explanations that are relevant to the text and coherent with human reasoning, with an average 48.76% better alignment with human judgment. |
Solving Math Word Problems via Cooperative Reasoning induced Language Models (2023.acl-long)
Copied to clipboard
Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, Yujiu Yang
| Challenge: | Large-scale pre-trained language models (PLMs) can be used to solve math word problems, but they lack fast adaptivity as humans. |
| Approach: | They propose a cooperative reasoning-induced PLM for solving the math word problem . they use system 1 as the generator and system 2 as the verifier to generate reasoning paths . |
| Outcome: | The proposed model improves on several mathematical reasoning datasets and achieves 9.6% improvement over baselines. |
XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for active learning rely on model uncertainty or disagreement to pick unlabeled data, leading to over-confidence in superficial patterns and lack of exploration. |
| Approach: | They propose to use a bi-directional encoder and a uni-directional decoder to generate and score an explanation for low-resource text classification. |
| Outcome: | The proposed model improves on 9 strong baselines on six datasets and can generate explanations for its predictions. |
Reverse Thinking Makes LLMs Stronger Reasoners (2025.naacl-long)
Copied to clipboard
Justin Chen, Zifeng Wang, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister
| Challenge: | Reverse-Enhanced Thinking (RevThink) is a framework for large language models to perform reverse thinking. |
| Approach: | They propose a framework for enhancing forward-backward reasoning by collecting data from a teacher model and employing three objectives to train a student model in a multi-task learning fashion. |
| Outcome: | The proposed framework outperforms a fine-tuning method trained on 10x more forward reasoning on 12 datasets covering commonsense, math, and logical reasoning. |
Human Rationales as Attribution Priors for Explainable Stance Detection (2021.emnlp-main)
Copied to clipboard
| Challenge: | In this work, we present a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. |
| Approach: | They propose a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. |
| Outcome: | The proposed method improves the reasoning of a state-of-the-art classifier in a data-scarce setting at no cost in predictive performance. |
Could you give me a hint ? Generating inference graphs for defeasible reasoning (2021.findings-acl)
Copied to clipboard
| Challenge: | Defeasible reasoning is a mode of reasoning where conclusions can be overturned by taking into account new evidence. |
| Approach: | They propose to automatically generate inference graphs for a defeasible inference task by transfer learning from a related NLP task. |
| Outcome: | The proposed method generates meaningful graphs for a defeasible inference task and human accuracy improves by 20%. |
Evaluating the Deductive Competence of Large Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing large language models have limited abilities to solve deductive reasoning problems . performance differences between conditions do not improve overall performance . |
| Approach: | They investigate whether several large language models can solve a deductive reasoning problem in their conventional form. |
| Outcome: | The proposed models can solve a classic type of deductive reasoning problem in their conventional form. |
IRAC: A Domain-Specific Annotated Corpus of Implicit Reasoning in Arguments (2022.lrec-1)
Copied to clipboard
| Challenge: | Using crowdsourcing, we show that models trained with domain-specific implicit reasonings outperform domain-general models in both automatic and human evaluations. |
| Approach: | They propose to create a domain-specific corpus of implicit reasonings annotated for a wide range of arguments and use it to generate models. |
| Outcome: | The proposed corpus outperforms domain-general models in automatic and human evaluations. |
Does Self-Rationalization Improve Robustness to Spurious Correlations? (2022.emnlp-main)
Copied to clipboard
| Challenge: | Rationalization is fundamental to human reasoning and learning. |
| Approach: | They evaluate robustness to spurious correlations in encoder-decoder and decoder-only models . authors say explanations can come at the cost of robustness . |
| Outcome: | The proposed model outputs are more interpretable and easier to interact with for end-users than nonrationalizing models. |
DetermLR: Augmenting LLM-based Logical Reasoning from Indeterminacy to Determinacy (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have revolutionized the landscape of reasoning tasks. |
| Approach: | They propose a new approach that rethinks the reasoning process as an evolution from indeterminacy to determinacy. |
| Outcome: | The proposed model surpasses all baselines on various logical reasoning benchmarks. |
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models (2025.coling-main)
Copied to clipboard
Kazuki Hayashi, Kazuma Onishi, Toma Suzuki, Yusuke Ide, Seiji Gobara, Shigeki Saito, Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe
| Challenge: | Large-scale vision language models excel at generating factual content, but their ability to rank images from multiple perspectives has not been explored. |
| Approach: | They propose a framework to evaluate large-scale vision-language models by measuring their ability to rank image texts from multiple perspectives. |
| Outcome: | The proposed evaluation framework measures how closely LVLMs' judgments align with human interpretations. |
StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding (2023.emnlp-main)
Copied to clipboard
Cheng Jiayang, Lin Qiu, Tsz Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, Zheng Zhang
| Challenge: | Analogy-making between narratives is crucial for human reasoning . despite its importance, there has been limited research on story analogies . |
| Approach: | They construct a large-scale story-level analogy corpus with 24K story pairs . they find that the tasks are incredibly difficult for large language models such as ChatGPT . |
| Outcome: | The proposed corpus contains 24K story pairs from diverse domains with human annotations on two similarities from the extended Structure-Mapping Theory. |
DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-level Value Function (2023.acl-long)
Copied to clipboard
Haiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, Jian Yin, Zhenguo Li, Xiaodan Liang
| Challenge: | Recent advances in neural theorem-proving resort to large language models and tree searches. |
| Approach: | They propose a Dynamic-Tree Driven Theorem Solver to accommodate general theoremes by guiding the search procedure with state confidence and proof-level values. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two popular theorem-proving datasets with a 6.65% improvement on average in terms of success rate. |
Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing concepts-based explainable approaches do not discover unseen concepts . a recent approach to solve this problem is concept-based explanations . |
| Approach: | They propose a framework that extracts comprehensible concepts automatically with no annotations . ECO-Concept uses an object-centric architecture to extract task-specific semantic concepts . |
| Outcome: | a new framework extracts comprehensible concepts with no concept annotations . the proposed framework outperforms existing methods in computability tests on diverse tasks . |
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion (2026.acl-long)
Copied to clipboard
| Challenge: | Existing quantization-based approaches to knowledge Graph Completion (KGC) are incomplete. |
| Approach: | They propose a framework that generates semantically coherent discrete codes for KG entities . they introduce a Granular Semantic Enhancement module that injects hierarchical knowledge into the codebook . |
| Outcome: | The proposed framework outperforms existing text-based and embedding-based baselines in the KGC domain. |
Semantic-Aware Logical Reasoning via a Semiotic Framework (2026.acl-long)
Copied to clipboard
Yunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang, Zikai Song
| Challenge: | Existing studies largely overlook the interplay between logical complexity and semantic complexity, limiting their robustness under abstract propositions, ambiguous contexts, and conflicting stances. |
| Approach: | They propose a semiotic-square-guided framework that integrates automated deduction with reflective verification to manage logical complexity across deeper reasoning chains. |
| Outcome: | The proposed framework achieves state-of-the-art performance on RepublicQA with 6.25% average gain, and generalizes well to four mainstream logical reasoning benchmarks with an additional 7.05% improvement. |
TORSO: Template-Oriented Reasoning Towards General Tasks (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to generate responses using few-shot examples depend on the provided examples, limiting the model’s reasoning capabilities. |
| Approach: | They propose a model that emulates human reasoning during response generation by using curated few-shot prompts instead of manually crafted few-shot examples. |
| Outcome: | The proposed model achieves strong performance on diverse LLMs benchmarks with reasonable rationales. |
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing models of large language reasoning exhibit reasoning biases similar to humans, a study shows . syllogistic reasoning is one of the basic forms of deductive reasoning . |
| Approach: | They propose to use a syllogism dataset to evaluate models' reasoning abilities . they propose to ask LLMs to translate slogismatic slurs into abstract logical expressions . |
| Outcome: | The proposed method shows that models exhibit reasoning biases similar to humans, and that there is room for improvement in reasoning problems where premises and hypotheses are neither entailment nor contradiction. |
IntentionESC: An Intention-Centered Framework for Enhancing Emotional Support in Dialogue Systems (2025.findings-acl)
Copied to clipboard
| Challenge: | IntentionESC defines the possible intentions of supporters in emotional support conversations, identifies key emotional state aspects for inferring these intentions, and maps them to appropriate support strategies. |
| Approach: | They propose an Intention-centered Emotional Support Conversation framework which defines the possible intentions of supporters in emotional support conversations, identifies key emotional state aspects for inferring intentions, and maps them to appropriate support strategies. |
| Outcome: | The proposed framework defines the possible intentions of supporters in emotional support conversations, identifies key emotional state aspects for inferring these intentions, and maps them to appropriate support strategies. |
Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy (2026.findings-acl)
Copied to clipboard
Vallabhaneni Raj Kumar, Ashwin S, Supriya Manna, Niladri Sett, Cheedella V S N M S Hema Harshitha, Kurakula Harshitha, Basina Deepakraj, Anand Kumar Sharma, Tanuj Sarkar, Samanthapudi Shakeer, Bondada Navaneeth Krishna
| Challenge: | a limited amount of annotated data has slowed progress in machine learning for low-resource languages . a sentiment label records an annotator's final decision, but it is not a valid record of the annotation's interpretation. |
| Approach: | They propose a large-scale Telugu sentiment classification dataset annotated with sentiment labels and human-selected rationales from multiple native speakers. |
| Outcome: | The proposed model improves classification performance, explanation quality, and social bias by incorporating human rationales. |
Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow and misaligned with human reasoning. |
| Approach: | They propose a value-grounded framework for evaluating and distilling structured moral reasoning in large language models. |
| Outcome: | The proposed framework evaluates 12 open-source models across four moral datasets. |
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset (2026.acl-long)
Copied to clipboard
YongKang Liu, Jiayang Yu, Mingyang Wang, Yiqun Zhang, Ercong Nie, Shi Feng, Daling Wang, Kaisong Song, Hinrich Schuetze
| Challenge: | Argumentation is a key part of human reasoning and decision-making . existing argumentative corpora focus on single-turn settings, but multi-turn dialogues are often realized as multi-turned dialogues . |
| Approach: | They present a dataset for strategic multi-turn argumentation dialogues . they annotate each utterance with five strategy types, allowing multiple strategies per utterrance . |
| Outcome: | The proposed dataset shows that explicit prompting improves fluency, stylistic coherence and persuasiveness. |
From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context (2026.acl-long)
Copied to clipboard
| Challenge: | Existing explanation methods for graph neural networks struggle to generate interpretable, fine-grained rationales. |
| Approach: | They propose a lightweight framework that uses large language models to generate interpretable explanations for GNNs. |
| Outcome: | The proposed framework generates interpretable explanations for GNN predictions using large language models. |
From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations (2026.findings-acl)
Copied to clipboard
| Challenge: | DeFactoX integrates Direct Preference Optimization (DPO) with Curriculum learning to align machine-generated explanations with human reasoning. |
| Approach: | They propose a framework that integrates Direct Preference Optimization (DPO) with Curriculum learning to align machine-generated explanations with human reasoning. |
| Outcome: | The proposed framework combines Direct Preference Optimization (DPO) with Curriculum learning to align machine-generated explanations with human reasoning. |